Papers with multilingual text detoxification

1 papers
SynthDetoxM: Modern LLMs are Few-Shot Parallel Detoxification Data Annotators (2025.naacl-long)

Copied to clipboard

Challenge: Existing approaches to multilingual text detoxification are hampered by the scarcity of parallel multilingual datasets.
Approach: They propose a pipeline for the generation of multilingual parallel detoxification data and a dataset for SynthDetoxM which is manually generated and rewritten with open-source LLMs.
Outcome: The proposed pipeline outperforms human-annotated datasets even in data limited setting.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations